Use checked rounding for BFC arena allocations - #32010
Merged
Akshay Sonawane (apsonawane) merged 1 commit intoAug 13, 2026
Merged
Conversation
Copilot started reviewing on behalf of
Akshay Sonawane (apsonawane)
August 12, 2026 05:42
View session
Akshay Sonawane (apsonawane)
enabled auto-merge (squash)
August 12, 2026 05:43
Contributor
There was a problem hiding this comment.
Pull request overview
This PR hardens BFCArena’s allocation-size rounding logic against integer overflow by switching the rounding arithmetic to SafeInt<size_t>, and adds a unit test to ensure overflow cases throw an OnnxRuntimeException instead of silently wrapping.
Changes:
- Updated
BFCArena::RoundedBytesto useSafeInt<size_t>for checked rounding arithmetic (preventing overflow onbytes + kMinAllocationSize - 1). - Added a regression test validating that an overflow-sized allocation request throws
OnnxRuntimeException. - Added required headers to support the new checked arithmetic and boundary-value test.
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.
| File | Description |
|---|---|
onnxruntime/core/framework/bfc_arena.cc |
Uses SafeInt<size_t> in BFCArena::RoundedBytes to make rounding overflow-safe. |
onnxruntime/test/framework/bfc_arena_test.cc |
Adds RoundedBytesOverflowThrows to assert overflow requests throw OnnxRuntimeException. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Ti-Tai Wang (titaiwangms)
approved these changes
Aug 13, 2026
Akshay Sonawane (apsonawane)
deleted the
fix/bfc-arena-rounded-bytes-overflow
branch
August 13, 2026 17:53
This was referenced Sep 10, 2026
Open
This was referenced Sep 19, 2026
Closed
This was referenced Sep 27, 2026
Kristoffer-Andre Kalliainen (181192)
pushed a commit
to 181192/brasscribe
that referenced
this pull request
Sep 30, 2026
Updated [AlphaSkia.Native.Linux](https://github.com/CoderLine/alphaSkia) from 3.4.135 to 3.5.147. <details> <summary>Release notes</summary> _Sourced from [AlphaSkia.Native.Linux's releases](https://github.com/CoderLine/alphaSkia/releases)._ ## 3.5.147 ## What's Changed * chore(deps): update skia and update patches by @Danielku15 in CoderLine/alphaSkia#69 **Full Changelog**: CoderLine/alphaSkia@v3.4.135...v3.5.147 Commits viewable in [compare view](CoderLine/alphaSkia@v3.4.135...v3.5.147). </details> Updated [AlphaSkia.Native.MacOs](https://github.com/CoderLine/alphaSkia) from 3.4.135 to 3.5.147. <details> <summary>Release notes</summary> _Sourced from [AlphaSkia.Native.MacOs's releases](https://github.com/CoderLine/alphaSkia/releases)._ ## 3.5.147 ## What's Changed * chore(deps): update skia and update patches by @Danielku15 in CoderLine/alphaSkia#69 **Full Changelog**: CoderLine/alphaSkia@v3.4.135...v3.5.147 Commits viewable in [compare view](CoderLine/alphaSkia@v3.4.135...v3.5.147). </details> Updated [Microsoft.ML.OnnxRuntime](https://github.com/Microsoft/onnxruntime) from 1.24.4 to 1.30.0. <details> <summary>Release notes</summary> _Sourced from [Microsoft.ML.OnnxRuntime's releases](https://github.com/Microsoft/onnxruntime/releases)._ ## 1.30.0 ONNX Runtime 1.30.0 expands generative AI inference, improves CPU and GPU performance, adds Go bindings, and strengthens runtime reliability. These notes cover changes since ONNX Runtime 1.29.1. ## Highlights - Expanded CUDA inference support with variable-length causal convolution for continuous batching, speculative decoding in paged XQA, and INT4 paged KV caches with per-channel scales ([#32168](microsoft/onnxruntime#32168), [#32340](microsoft/onnxruntime#32340), [#32515](microsoft/onnxruntime#32515)). - Improved WebGPU PagedAttention, added GPT-OSS support and INT8 KV-cache block quantization, and extended convolution optimizations ([#31727](microsoft/onnxruntime#31727), [#32277](microsoft/onnxruntime#32277), [#32284](microsoft/onnxruntime#32284), [#32420](microsoft/onnxruntime#32420)). - Added fused CPU LinearAttention kernels for AVX-512, Arm64 NEON, and SVE, plus AVX2 LayerNorm/RMSNorm acceleration ([#31674](microsoft/onnxruntime#31674), [#31973](microsoft/onnxruntime#31973), [#32178](microsoft/onnxruntime#32178), [#32356](microsoft/onnxruntime#32356)). - Added Go bindings for the ONNX Runtime C API and DeepSeek Engram contrib operators ([#29615](microsoft/onnxruntime#29615), [#32268](microsoft/onnxruntime#32268)). ## Announcements & Compatibility - FP4 QMoE kernels are now enabled by default in CUDA builds, with Windows build support added in this release. Source builds can opt out with `-Donnxruntime_USE_FP4_QMOE=OFF` ([#32096](microsoft/onnxruntime#32096), [#32163](microsoft/onnxruntime#32163)). - CUDA fpA-intB builds now default to a compact kernel set for FP16 activations, INT4/INT8 weights, scale-only quantization, and `block_size=32`. Set `-Donnxruntime_USE_FPA_INTB_GEMM_FULL=ON` when building from source to retain the full kernel set, including BF16, zero-point, bias, larger-block-size, and native Hopper variants ([#32324](microsoft/onnxruntime#32324)). - CPU FP16 `Gemm` and `MatMul` execution is gated on hardware acceleration. CPU-assigned FP16 nodes without a matching kernel now fall back to FP32 ([#32301](microsoft/onnxruntime#32301), [#32197](microsoft/onnxruntime#32197)). - WebGPU plugin EP packaging now supports Linux AArch64. Plugin versions were advanced to WebGPU 0.4.0 and CUDA 0.2 ([#32287](microsoft/onnxruntime#32287), [#31960](microsoft/onnxruntime#31960), [#31970](microsoft/onnxruntime#31970)). ## Security & Reliability ### Model Loading, Memory, and Input Validation - Limited nested model-graph depth and canonicalized external-data locations to harden model loading ([#32344](microsoft/onnxruntime#32344), [#32135](microsoft/onnxruntime#32135)). - Added checked rounding for BFC arena allocations and fixed prepacked-weight reference lifetimes ([#32010](microsoft/onnxruntime#32010), [#32040](microsoft/onnxruntime#32040)). - Strengthened shape, rank, and parameter validation for `Split`, `Scan`, `GatherND`, `ScatterND`, `SpaceToDepth`/`DepthToSpace`, `Crop`, `Conv`, `Normalizer`, and pooling ([#29461](microsoft/onnxruntime#29461), [#31668](microsoft/onnxruntime#31668), [#32034](microsoft/onnxruntime#32034), [#32039](microsoft/onnxruntime#32039), [#32076](microsoft/onnxruntime#32076), [#32157](microsoft/onnxruntime#32157), [#32160](microsoft/onnxruntime#32160), [#32161](microsoft/onnxruntime#32161), [#32345](microsoft/onnxruntime#32345), [#32349](microsoft/onnxruntime#32349)). - Hardened generation and attention input handling, including attention-attribute narrowing, `BifurcationDetector` inputs, generation subgraph shapes, and QEmbed segment inputs. BeamSearch buffer expansion now uses dynamic shape storage ([#31648](microsoft/onnxruntime#31648), [#31701](microsoft/onnxruntime#31701), [#32009](microsoft/onnxruntime#32009), [#32078](microsoft/onnxruntime#32078), [#32144](microsoft/onnxruntime#32144)). - Validated `TreeEnsemble` node references and bounded subtree comparison, rejected non-finite CPU `RoiAlign` coordinates, and required `ImageScaler` bias to match the channel count ([#32031](microsoft/onnxruntime#32031), [#32043](microsoft/onnxruntime#32043), [#32011](microsoft/onnxruntime#32011), [#32002](microsoft/onnxruntime#32002)). - Added an allowlist of safe LoRA adapter parameter data types, validated `MatMulFpQ4` shape inputs, and checked MLAS blockwise quantization/dequantization index ranges ([#31682](microsoft/onnxruntime#31682), [#32032](microsoft/onnxruntime#32032), [#32007](microsoft/onnxruntime#32007)). ### GPU Bounds and Resource Lifetimes - Hardened CUDA indexing and buffer-size arithmetic in `MatMulNBits`, `RemovePadding`, `RotaryEmbedding`, `SparseAttention`, Whisper beam search, NMS, QDQ, and `GatherElements` ([#31643](microsoft/onnxruntime#31643), [#31994](microsoft/onnxruntime#31994), [#31995](microsoft/onnxruntime#31995), [#31996](microsoft/onnxruntime#31996), [#31998](microsoft/onnxruntime#31998), [#32014](microsoft/onnxruntime#32014), [#32029](microsoft/onnxruntime#32029), [#32030](microsoft/onnxruntime#32030)). - Fixed overflow in CUDA reduction scans and Softmax offset arithmetic, and handled zero-sized outputs in CUDA random-generator kernels ([#32137](microsoft/onnxruntime#32137), [#32330](microsoft/onnxruntime#32330), [#31997](microsoft/onnxruntime#31997)). - Fixed CUDA MultiHeadAttention shared-cache scratch lifetimes and kept `CudaAsyncBuffer` staging storage alive across CUDA graph replay ([#31968](microsoft/onnxruntime#31968), [#32121](microsoft/onnxruntime#32121)). - Fixed WebGPU out-of-bounds subgroup-matrix loads for partial tiles, zero-initialized writable device-allocator buffers, and rejected foreign GPU handles in built-in data transfers ([#32364](microsoft/onnxruntime#32364), [#32063](microsoft/onnxruntime#32063), [#32317](microsoft/onnxruntime#32317)). ### Dependencies and Tooling - Upgraded Protobuf to 33.6 and refreshed Python documentation dependencies, including an ONNX security-related update ([#29906](microsoft/onnxruntime#29906), [#32190](microsoft/onnxruntime#32190), [#32424](microsoft/onnxruntime#32424)). - Updated JavaScript dependencies including `js-yaml`, `joi`, `fast-uri`, and the Next.js end-to-end fixture ([#32397](microsoft/onnxruntime#32397), [#32486](microsoft/onnxruntime#32486), [#32488](microsoft/onnxruntime#32488), [#32505](microsoft/onnxruntime#32505), [#32508](microsoft/onnxruntime#32508)). - Pinned GitHub Actions to full-length commit SHAs and strengthened packaging infrastructure with authenticated package feeds and NPM network isolation ([#32176](microsoft/onnxruntime#32176), [#32005](microsoft/onnxruntime#32005), [#32440](microsoft/onnxruntime#32440)). ## New Features ### Core APIs & Runtime - Added Go bindings for the ONNX Runtime C API ([#29615](microsoft/onnxruntime#29615)). - Extended memory importing with host-pointer support and added access to preallocated outputs through `KernelContext::GetPreallocatedOutput` ([#29726](microsoft/onnxruntime#29726), [#32089](microsoft/onnxruntime#32089)). - Added packed-attention workspace recipes and estimates, and made workspace input-shape handling aware of optional inputs ([#32283](microsoft/onnxruntime#32283), [#32321](microsoft/onnxruntime#32321), [#32312](microsoft/onnxruntime#32312)). - Added DeepSeek Engram contrib operators, `EngramGate` and `NGramHashMapping`, and expanded kernel coverage for Qwen-3.5 operators ([#32268](microsoft/onnxruntime#32268), [#32106](microsoft/onnxruntime#32106)). ### Plugin Execution Providers ... (truncated) ## 1.29.1 This is a patch release on top of [v1.29.0](https://github.com/microsoft/onnxruntime/releases/tag/v1.29.0), containing GroupQueryAttention capability and KV-cache layout improvements, plugin Execution Provider performance tooling updates, and targeted graph and optimizer fixes. ## GroupQueryAttention - Added bidirectional GroupQueryAttention support on CPU and CUDA through a backward-compatible `causal` attribute, with explicit handling for unsupported execution paths ([#31704](microsoft/onnxruntime#31704)) - Added a session option and Execution Provider metadata contract for using the BNHS Value KV-cache layout, with graph transformations that preserve compatibility with the existing BNSH operator schema ([#32139](microsoft/onnxruntime#32139)) - Added CPU support for `attention_bias` with a sliding-window KV cache, including explicit position IDs and post-eviction bias indexing ([#32302](microsoft/onnxruntime#32302)) ## Runtime and Performance Tools - Fixed Compile API model serialization when output-model and custom initializer-location callbacks are used together, preventing duplicate graph fields in emitted models ([#32303](microsoft/onnxruntime#32303)) - Updated `onnxruntime_perf_test` to use plugin Execution Provider device allocators for generated inputs, loaded test data, and pre-allocated outputs, avoiding unnecessary per-run host/device copies ([#32244](microsoft/onnxruntime#32244)) ## Bug Fixes and Documentation - Hardened FastGelu fusion to skip malformed `Mul` and `Pow` patterns ([#32016](microsoft/onnxruntime#32016)) - Added validation for in-memory external initializer references, rejecting unregistered or mismatched data before graph transformation ([#32042](microsoft/onnxruntime#32042)) - Restored the C API documentation workflow by switching the pinned Doxygen download to the official GitHub release asset ([#32210](microsoft/onnxruntime#32210)) ## Contributors Thanks to our 7 contributors for this release! [@adrastogi](https://github.com/adrastogi), [@apsonawane](https://github.com/apsonawane), [@edgchen1](https://github.com/edgchen1), [@javier-intel](https://github.com/javier-intel), [@jnagi-intel](https://github.com/jnagi-intel), [@tianleiwu](https://github.com/tianleiwu), [@Wayne-Ch](https://github.com/Wayne-Ch) <sub>Release highlights were drafted with AI assistance and are subject to release-team review.</sub> Full Changelog: [v1.29.0...v1.29.1](microsoft/onnxruntime@v1.29.0...v1.29.1) ## 1.29.0 ## Announcements & Breaking Changes - onnxruntime-web has announced the deprecation of WebGL and JSEP. The native WebGPU EP is the recommended path going forward. See the deprecation and migration plans for details ([#29716](microsoft/onnxruntime#29716), [#31683](microsoft/onnxruntime#31683)). - POSIX telemetry is now available on Linux, macOS, Android, and iOS when ONNX Runtime is built with telemetry enabled. It does not change the public ABI, WebAssembly remains telemetry-free, and setting `ORT_DISABLE_TELEMETRY=1` before initialization disables non-Windows telemetry for the process ([#27379](microsoft/onnxruntime#27379), [#29872](microsoft/onnxruntime#29872)). - The unused internal `onnxruntime/python/tools/tensorrt` dashboard tooling was removed. This does not affect the TensorRT Execution Provider APIs ([#29395](microsoft/onnxruntime#29395)). ## Security Fixes ### Path, bounds, and input validation - Fixed a path traversal vulnerability in TensorRT and NvTensorRTRTX engine refitting by making external-data path validation unconditional ([#29396](microsoft/onnxruntime#29396)). - Validated the CPU MoE `k` attribute against the number of experts and fixed a CPU `TensorScatter` security issue ([#29907](microsoft/onnxruntime#29907), [#29916](microsoft/onnxruntime#29916)). - Added missing rank, shape, and parameter validation for pooling, LSTM and DynamicQuantizeLSTM, Sampling, FeatureVectorizer, SkipLayerNorm, QLinearConv, Whisper decoding, RNN activations, GridSample, contrib `Range`, and `CropAndResize` ([#29254](microsoft/onnxruntime#29254), [#29255](microsoft/onnxruntime#29255), [#29265](microsoft/onnxruntime#29265), [#29579](microsoft/onnxruntime#29579), [#29595](microsoft/onnxruntime#29595), [#29605](microsoft/onnxruntime#29605), [#29871](microsoft/onnxruntime#29871), [#31636](microsoft/onnxruntime#31636), [#31671](microsoft/onnxruntime#31671), [#31675](microsoft/onnxruntime#31675), [#31676](microsoft/onnxruntime#31676), [#31684](microsoft/onnxruntime#31684)). - Hardened CUDA indexing and buffer handling in GridSample, transpose, GatherBlockQuantized, InstanceNormalization, LayerNorm/RMSNorm, BeamSearch, DeformConv, AveragePool, and MaxPool ([#29581](microsoft/onnxruntime#29581), [#29631](microsoft/onnxruntime#29631), [#29638](microsoft/onnxruntime#29638), [#31640](microsoft/onnxruntime#31640), [#31642](microsoft/onnxruntime#31642), [#31644](microsoft/onnxruntime#31644), [#31645](microsoft/onnxruntime#31645), [#31647](microsoft/onnxruntime#31647), [#31650](microsoft/onnxruntime#31650)). - Fixed packed sub-byte tensor over-copying in `OrtApi::GetValue` and validated DML constant tensor byte sizes ([#29157](microsoft/onnxruntime#29157), [#31665](microsoft/onnxruntime#31665)). ### Supply chain and tooling - Updated npm lockfiles, refreshed the Next.js end-to-end fixture lockfile for security advisories, and upgraded `adm-zip` for `onnxruntime-node` ([#29827](microsoft/onnxruntime#29827), [#29926](microsoft/onnxruntime#29926), [#31192](microsoft/onnxruntime#31192)). ## New Features ### Core APIs & Runtime - Default intra-op and inter-op thread-pool sizes can now be set with `ORT_INTRA_OP_NUM_THREADS` and `ORT_INTER_OP_NUM_THREADS`. Explicit thread settings still take precedence, and `0` preserves machine-sized defaults ([#29688](microsoft/onnxruntime#29688)). - Added weightless-model support for all initializer types, allowed zero-input `EpContext` nodes, and wired maximum-shape inference into workspace estimation ([#29607](microsoft/onnxruntime#29607), [#29799](microsoft/onnxruntime#29799), [#31613](microsoft/onnxruntime#31613)). - Added ONNX-domain support for rotary embedding and a fused `MRotaryEmbedding` contrib operator for Qwen mRoPE variants ([#29261](microsoft/onnxruntime#29261), [#31728](microsoft/onnxruntime#31728)). - Added multi-shape profiling to `onnxruntime_perf_test` through `--data_shape`, plus verbose graph-transformer tracing and broader inference-session error-path coverage ([#29555](microsoft/onnxruntime#29555), [#29558](microsoft/onnxruntime#29558), [#29569](microsoft/onnxruntime#29569), [#29571](microsoft/onnxruntime#29571)). ### Execution Provider ABI & Plugin EPs - WebGPU now supports device-free compile-only sessions for offline graph transformation ([#29681](microsoft/onnxruntime#29681)). - Expanded CUDA plugin EP packaging and testing, including Windows ARM64 package and size options, updated package outputs, and aligned architecture selections across Python, C API, TensorRT, Node.js, and plugin packages ([#31635](microsoft/onnxruntime#31635), [#31722](microsoft/onnxruntime#31722), [#31992](microsoft/onnxruntime#31992)). - Improved plugin lifecycle handling by unloading failed EP library loads and fixing allocator-deleter lifetime ([#29634](microsoft/onnxruntime#29634), [#29770](microsoft/onnxruntime#29770)). ## Execution Provider Updates ### NVIDIA CUDA EP #### Attention and decoding - Added `PagedAttention` with quantized KV cache, XQA decode, MLA, QK-Norm, and head-sink support ([#29912](microsoft/onnxruntime#29912)). - Extended quantized KV-cache support with attention sinks, independent and per-channel scales, sliding-window cache support, and a fused K/V dequantization launch ([#29900](microsoft/onnxruntime#29900), [#29904](microsoft/onnxruntime#29904), [#31480](microsoft/onnxruntime#31480)). - Added a cuDNN SDPA decode tier to the standard ONNX `Attention` CUDA kernel and enabled cuDNN SDPA for contrib `Attention` ([#29715](microsoft/onnxruntime#29715), [#29717](microsoft/onnxruntime#29717)). - Added `attention_bias` support to the GroupQueryAttention unfused path and `state_window` support to LinearAttention and CausalConvWithState for MTP ([#29525](microsoft/onnxruntime#29525), [#31157](microsoft/onnxruntime#31157)). - Fixed LinearAttention on GPUs with limited shared memory ([#31982](microsoft/onnxruntime#31982)). #### MoE and quantized GEMM - Added NVFP4 QMoE, including native FP4xFP4 prefill on SM120, faster decode GEMV, fused routing/finalization paths, and reduced activation and weight-dequantization overhead ([#29697](microsoft/onnxruntime#29697), [#29824](microsoft/onnxruntime#29824), [#29887](microsoft/onnxruntime#29887), [#29919](microsoft/onnxruntime#29919), [#31156](microsoft/onnxruntime#31156), [#31159](microsoft/onnxruntime#31159), [#31349](microsoft/onnxruntime#31349), [#31479](microsoft/onnxruntime#31479)). ... (truncated) ## 1.28.2 This is a patch release on top of [v1.28.1](https://github.com/microsoft/onnxruntime/releases/tag/v1.28.1), containing a targeted fix for Compile API model serialization. ## Highlights ### Bug Fixes - Fixed Compile API callback serialization to prevent duplicate graph nodes, inputs, outputs, and value information in emitted optimized models, including models with embedded or external initializers ([#32303](microsoft/onnxruntime#32303)) ## Contributors Thanks to our contributor for this release! [@adrastogi](https://github.com/adrastogi) Full Changelog: [v1.28.1...v1.28.2](microsoft/onnxruntime@v1.28.1...v1.28.2) > These release notes were drafted with assistance from GitHub Copilot. ## 1.28.1 This is a patch release on top of [v1.28.0](https://github.com/microsoft/onnxruntime/releases/tag/v1.28.0), containing support for device-free WebGPU compilation, improved compatibility with sandboxed Windows processes, and targeted graph-validation fixes. ## WebGPU EP - Added support for device-free compile-only sessions, enabling offline graph transformation and optimized-model serialization without access to GPU hardware ([#29681](microsoft/onnxruntime#29681)) ## Bug Fixes - Prevented an access violation in Windows processes under Win32k lockdown by skipping DXGI device discovery ([#29755](microsoft/onnxruntime#29755)) - Allowed zero-input `EPContext` nodes, aligning their schema with support for compiling zero-input models ([#29799](microsoft/onnxruntime#29799)) - Hardened FastGelu fusion to skip malformed `Mul` and `Pow` patterns ([#32016](microsoft/onnxruntime#32016)) - Added validation for in-memory external initializer references, rejecting unregistered or mismatched data before graph transformation ([#32042](microsoft/onnxruntime#32042)) ## Contributors Thanks to our 4 contributors for this release! [@apsonawane](https://github.com/apsonawane), [@shiyi9801](https://github.com/shiyi9801), [@adrastogi](https://github.com/adrastogi), [@mingmingtasd](https://github.com/mingmingtasd) Full Changelog: [v1.28.0...v1.28.1](microsoft/onnxruntime@v1.28.0...v1.28.1) ## 1.28.0 ## Announcements & Breaking Changes - Upgraded to **ONNX 1.22.0** and protobuf 6.33.5 ([#28754](microsoft/onnxruntime#28754), [#29606](microsoft/onnxruntime#29606), [#28967](microsoft/onnxruntime#28967)). Graph optimizer opset version checks were updated accordingly ([#28966](microsoft/onnxruntime#28966)). - **cuDNN and cuFFT are now optional at runtime** for the CUDA EP, and `nvrtc` is no longer linked, which significantly reduces the required CUDA redistributable footprint ([#29252](microsoft/onnxruntime#29252), [#29808](microsoft/onnxruntime#29808), [#29705](microsoft/onnxruntime#29705), [#29620](microsoft/onnxruntime#29620)). - An **experimental C/C++ API surface** was introduced. `OrtModelPackageApi` now lives in the experimental C API and may change in future releases ([#28746](microsoft/onnxruntime#28746), [#29142](microsoft/onnxruntime#29142), [#28990](microsoft/onnxruntime#28990)). - **Deprecated / removed:** - SkipLayerNorm strict mode is deprecated ([#29388](microsoft/onnxruntime#29388)). - The TensorRT fused causal attention kernels were removed from the CUDA EP ([#29143](microsoft/onnxruntime#29143)). - The dynamic WGSL generator (duktape/Node) path was removed in favor of the Python `wgsl-gen` implementation ([#29141](microsoft/onnxruntime#29141), [#28355](microsoft/onnxruntime#28355)). - `CUDA_QUANT_PREPROCESS` is off by default ([#29687](microsoft/onnxruntime#29687)). - NPM packages are now published from the CUDA 13 pipeline ([#28773](microsoft/onnxruntime#28773)). - The CUDA 12.8 package architecture list was refreshed for this release ([#29711](microsoft/onnxruntime#29711)). ## Security Fixes ### Memory safety & input validation - Hardened the ORT FlatBuffer model loader against malformed buffers, and removed now-redundant table offset validation ([#28186](microsoft/onnxruntime#28186), [#29068](microsoft/onnxruntime#29068)) - Fixed type confusion in raw-pointer `bind_input` causing an out-of-bounds write ([#28839](microsoft/onnxruntime#28839)) - Fixed out-of-bounds pointer in `TensorAt` for sub-byte packed types ([#28973](microsoft/onnxruntime#28973)) - Fixed arbitrary memory read, out-of-bounds dereference, and other OOB accesses in kernels ([#28991](microsoft/onnxruntime#28991), [#29011](microsoft/onnxruntime#29011), [#29012](microsoft/onnxruntime#29012), [#29014](microsoft/onnxruntime#29014)) - Validated `Col2Im` inputs to prevent heap over-read ([#28706](microsoft/onnxruntime#28706)) - Hardened `CropAndResize` against malformed `crop_size` tensors ([#28766](microsoft/onnxruntime#28766)) - Validated `BeamSearch` `vocab_size` against logits width ([#28774](microsoft/onnxruntime#28774)) - Fixed bounds in `WhisperDecoderSubgraph::CreateInitialFeeds` ([#29239](microsoft/onnxruntime#29239)) - Validated `SparseAttention` CSR indices/key lengths and rejected zero-dimension `block_row_indices` ([#29015](microsoft/onnxruntime#29015), [#29242](microsoft/onnxruntime#29242)) - Clamped derived sequence lengths and KV-cache index in CUDA GroupQueryAttention, and fixed a CPU GQA out-of-bounds read in the past-KV buffer ([#29240](microsoft/onnxruntime#29240), [#29447](microsoft/onnxruntime#29447)) - Clamped 1D attention `mask_index` to valid bounds ([#29449](microsoft/onnxruntime#29449)) - Validated `MaxpoolWithMask` kernel rank against input spatial rank ([#29253](microsoft/onnxruntime#29253)) - Rejected CUDA BERT `EmbedLayerNorm`/`SkipLayerNorm` shapes exceeding 32-bit output indexing ([#29264](microsoft/onnxruntime#29264)) - Fixed the optional-output guard in `DecoderAttention`/`MultiHeadAttention` shape inference and negative-axis handling in `ExpandDims` shape inference ([#29268](microsoft/onnxruntime#29268), [#29448](microsoft/onnxruntime#29448)) - Fixed `TreeEnsemble` target id validation and added input validation to `LinearClassifier` ([#29293](microsoft/onnxruntime#29293), [#29060](microsoft/onnxruntime#29060)) - Fixed `DynamicQuantizeLSTM` zero-point/scale validation typos ([#29462](microsoft/onnxruntime#29462)) - Handled non-trivially-copyable types in `Loop`/`Scan` output concatenation ([#29397](microsoft/onnxruntime#29397)) - Normalized bool tensor `raw_data` to `{0, 1}` on unpack ([#29238](microsoft/onnxruntime#29238)) - Addressed hardening gaps in `Resize`, `PadFusion`, and LoRA handling ([#28779](microsoft/onnxruntime#28779), [#28780](microsoft/onnxruntime#28780), [#28801](microsoft/onnxruntime#28801)) - Fixed unbounded lifetime on `WithOutputTensor` in the Rust bindings ([#29251](microsoft/onnxruntime#29251)) ### Integer overflow & allocation size - Guarded `MlasConvPrepare` working-buffer products and `ConvTranspose` pad computation with SafeInt ([#29444](microsoft/onnxruntime#29444), [#29446](microsoft/onnxruntime#29446)) - Fixed signed-int overflow in `SamplingState::Init` that could cause a heap buffer overflow ([#29443](microsoft/onnxruntime#29443)) - Hardened QMoE against integer overflow and partial K tiles ([#29067](microsoft/onnxruntime#29067)) - Validated `B`/scales/zero-points shape in `MatMulNBits::PrePack` ([#29445](microsoft/onnxruntime#29445)) - Pre-checked `ConstantOfShape` output size against the input initializer before constant folding ([#28751](microsoft/onnxruntime#28751)) - Fixed integer overflow in RKNPU implicit bias allocation ([#29249](microsoft/onnxruntime#29249)) - Fixed WebGPU out-of-bounds reads in `Pad` (int64/int32 truncation), `Slice`, and `GatherBlockQuantized` ([#28721](microsoft/onnxruntime#28721), [#28704](microsoft/onnxruntime#28704), [#28718](microsoft/onnxruntime#28718)) ### Supply chain & tooling ... (truncated) ## 1.27.1 This is a patch release on top of [v1.27.0](https://github.com/microsoft/onnxruntime/releases/tag/v1.27.0), containing targeted bug fixes, a CUDA QMoE decode-path optimization, and CI/build infrastructure fixes. ## Bug Fixes - [MLAS] Fixed an `igemm` regression in the KleidiAI path ([#28571](microsoft/onnxruntime#28571)) - Fixed a QMoE CPU livelock by eliminating nested intra-op parallelism ([#29081](microsoft/onnxruntime#29081)) - Fixed a regression in graph-capture session initialization that rejected an empty graph ([#29457](microsoft/onnxruntime#29457)) - Fixed CustomOp forward compatibility by capping the version instead of rejecting it ([#29574](microsoft/onnxruntime#29574)) ## Performance ### NVIDIA CUDA EP - Added a QMoE GEMV fast path for batch-1 decode ([#29038](microsoft/onnxruntime#29038)) ## CI & Build Infrastructure - Fixed an incorrect identity for `azcopy` ([#29274](microsoft/onnxruntime#29274)) - Fixed a `brew install applesimutils` failure by trusting the wix/brew tap ([#29450](microsoft/onnxruntime#29450)) - Upgraded to Xcode 26 ([#29468](microsoft/onnxruntime#29468)) - Stopped echoing the command when setting a VSO variable in `mac-cpu-packing-jobs.yml` ([#29575](microsoft/onnxruntime#29575)) - Fixed the web e2e (npm/vite) and Python DML CI pipelines ([#29609](microsoft/onnxruntime#29609)) ## Contributors Thanks to our 8 contributors for this release! [@tianleiwu](https://github.com/tianleiwu), [@chilo-ms](https://github.com/chilo-ms), [@edgchen1](https://github.com/edgchen1), [@adrastogi](https://github.com/adrastogi), [@damdoo01-arm](https://github.com/damdoo01-arm), [@JonathanC-ARM](https://github.com/JonathanC-ARM), [@martin-klacer-arm](https://github.com/martin-klacer-arm), [@sanaa-hamel-microsoft](https://github.com/sanaa-hamel-microsoft) Full Changelog: [v1.27.0...v1.27.1](microsoft/onnxruntime@v1.27.0...v1.27.1) ## 1.27.0 n.b. This release is targeting ONNX 1.21. ONNX 1.22 will be supported in ORT 1.28. n.b. This changelog was generated via LLM. Only the contributor list has been verified. As always, only trust the commit history. ## Announcements & Breaking Changes * CUDA 12 package files are now explicitly named as such. * CUDA 12 packages are deprecated, please move to CUDA 13 ASAP. --- ## Security Fixes * Fixed out-of-bounds read in `SoftmaxCrossEntropyLoss` via label bounds validation ([#28004](microsoft/onnxruntime#28004)) * Hardened `OneHot` input validation and output-size computation ([#28014](microsoft/onnxruntime#28014)) * Added SafeInt overflow protection in `Expand` and capped constant-folding output sizes ([#28055](microsoft/onnxruntime#28055)) * Bounded total output allocation size in `Tile` kernel ([#28070](microsoft/onnxruntime#28070)) * Added mask/input shape consistency checks in `MaxpoolWithMask::Compute` ([#28223](microsoft/onnxruntime#28223)) * Fixed `BitShift` UB for shift amounts greater than or equal to bit width ([#28272](microsoft/onnxruntime#28272)) * Validated sequence bounds in GQA (`seqlens_k` vs `cos_cache`) ([#28277](microsoft/onnxruntime#28277)) * Validated conv bias shape in `WordConvEmbedding` to prevent OOB reads ([#28279](microsoft/onnxruntime#28279)) * Fixed int32 overflow in CUDA Cast and UnaryElementWise kernels for very large tensors ([#28386](microsoft/onnxruntime#28386)) * Fixed out-of-bounds read in `CropBase` scale handling ([#28399](microsoft/onnxruntime#28399)) * Fixed rank-underflow bug in Inverse kernel trailing-dimension indexing ([#28400](microsoft/onnxruntime#28400)) * Added sparse tensor external file path validation and additional external-path hardening ([#28408](microsoft/onnxruntime#28408), [#28709](microsoft/onnxruntime#28709), [#28725](microsoft/onnxruntime#28725)) * Switched remaining `torch.load()` calls to `weights_only=True` ([#28421](microsoft/onnxruntime#28421)) * Added CPU cache-indirection beam-index validation ([#28486](microsoft/onnxruntime#28486)) * Added additional overflow/bounds checks and test coverage in runtime buffers ([#28713](microsoft/onnxruntime#28713), [#28747](microsoft/onnxruntime#28747)) --- ## New Features ### Execution Provider Plugin API * Added zero-copy I/O for plugin EPs with HOST_ACCESSIBLE memory ([#28037](microsoft/onnxruntime#28037)) * Added `OrtEp::OnSessionInitializationEnd()` callback ([#28319](microsoft/onnxruntime#28319)) * Added plugin EP session-options getters ([#28377](microsoft/onnxruntime#28377)) * Added CUDA Plugin EP provider options for streams and external allocators ([#28603](microsoft/onnxruntime#28603)) ### Core APIs & Runtime * Added support for ONNX overloaded functions (IR v10+) ([#28275](microsoft/onnxruntime#28275)) * Added FLOAT8E8M0 datatype support in ONNX Runtime ([#28381](microsoft/onnxruntime#28381)) * Added CPU Cast support for FLOAT8E8M0 ([#28435](microsoft/onnxruntime#28435)) * Added `kOrtEpDevice_EpMetadataKey_OSDriverVersion` example and docs ([#28282](microsoft/onnxruntime#28282)) ### Quantization & Training Tooling * Added calibration cache support to `quantize_static` ([#28221](microsoft/onnxruntime#28221)) * Added `ActivationRestrictedAsymmetric` quantization option ([#28237](microsoft/onnxruntime#28237)) ... (truncated) ## 1.26.0 n.b. The following was generated via LLM from Git history. Only the contributor list has been verified. ## Announcement - Breaking Changes - **Support for CUDA 12 will be removed in 1.27.0.** - CUDA 13 will continue to be published as `onnxruntime-<os>-<arch>-gpu_cuda13-<version>.<ext>` - CUDA runtime will be moving soon to a dedicated Execution Provider (EP) instead of a published package from ORT core. ## Highlights - Added optional memory mapping for `.ort` model loads ([#28164](microsoft/onnxruntime#28164)). - Added RISC-V Vector (RVV) support for CPU EP ([#28261](microsoft/onnxruntime#28261)). - OpenVINO EP upgraded for 1.26.0 development release ([#28297](microsoft/onnxruntime#28297)). - WebGPU gained GridSample support ([#28264](microsoft/onnxruntime#28264)) and Split-K improvements ([#28151](microsoft/onnxruntime#28151)). - CUDA plugin EP gained graph support ([#28002](microsoft/onnxruntime#28002)), profiling API ([#28216](microsoft/onnxruntime#28216)). ## Security and Reliability Hardening - Replaced unrestricted Python `setattr` configuration with an allowlist ([#28083](microsoft/onnxruntime#28083)). - Hardened multiple OOB and overflow scenarios across ML and core ops: - Attention mask index OOB write ([#27789](microsoft/onnxruntime#27789)). - MaxPoolGrad indices bounds validation ([#27903](microsoft/onnxruntime#27903)). - SVM and TreeEnsemble bounds/security fixes ([#27950](microsoft/onnxruntime#27950), [#27951](microsoft/onnxruntime#27951), [#27952](microsoft/onnxruntime#27952), [#27989](microsoft/onnxruntime#27989)). - RNN sequence_lens OOB read and integer overflow handling ([#28052](microsoft/onnxruntime#28052), [#28003](microsoft/onnxruntime#28003)). - GroupQueryAttention seqlens_k bounds validation and compatibility follow-up ([#28031](microsoft/onnxruntime#28031), [#28259](microsoft/onnxruntime#28259)). - MatMulBnb4 and ML coefficient SafeInt checks ([#27995](microsoft/onnxruntime#27995), [#28001](microsoft/onnxruntime#28001)). - CUDA Gather int32 overflow fix ([#28108](microsoft/onnxruntime#28108)). - GridSample float->int64 cast hardening for NaN/Inf/out-of-range coords ([#28302](microsoft/onnxruntime#28302)). - Fixed session logger use-after-free during EP teardown under verbose logging ([#28274](microsoft/onnxruntime#28274)). ## CUDA, Attention, and MLAS - Filled CUDA opset/operator gaps and extended support: - Transpose opset 23 -> 25 ([#27740](microsoft/onnxruntime#27740)). - QuantizeLinear/DequantizeLinear opset 25 ([#28046](microsoft/onnxruntime#28046)). - CUDA TopK INT8/INT16/UINT8 support ([#27862](microsoft/onnxruntime#27862)). - LabelEncoder CUDA support for numeric types ([#28045](microsoft/onnxruntime#28045)). - Attention/GQA improvements: - Fixed ONNX Attention min-bias alignment crash on SM<80 and masked-batch NaN behavior ([#27831](microsoft/onnxruntime#27831)). - Added FP32 QK accumulation path for unfused GQA attention ([#28198](microsoft/onnxruntime#28198)). - Added CUDART_VERSION reduction compatibility in GQA attention ([#28296](microsoft/onnxruntime#28296)). - Fixed CUDA 13 build error in GQA unfused attention ([#28309](microsoft/onnxruntime#28309)). - PagedAttention fallback for SM<80 fp16 ([#28200](microsoft/onnxruntime#28200)). - MLAS updates: - FP16 Gelu enablement ([#26815](microsoft/onnxruntime#26815)). - Arm64 BF16 fast-math conv kernels for NCHW/NCHWc paths ([#27878](microsoft/onnxruntime#27878)). ## WebGPU, WebNN, and JavaScript - WebGPU feature and correctness updates: ... (truncated) ## 1.25.1 n.b. This changelog is LLM generated. Only the contributor listing has been verified. # ONNX Runtime Release 1.25.1 ## 📢 Announcements & Breaking Changes ### ONNX Op Updates * **Enhanced ONNX operator support** with new opset versions: Reshape (opset 25), Transpose (opset 24) ([#27752](microsoft/onnxruntime#27752)) --- ## ✨ New Features ### 📊 New ONNX Ops & Model Support * **LinearAttention and CausalConvState operators** for Qwen3.5 model support ([#27907](microsoft/onnxruntime#27907)) * **RotaryEmbedding (RotEMB) and RMSNorm operators** added ([#27752](microsoft/onnxruntime#27752)) * **Linear Attention signature** support ([#27842](microsoft/onnxruntime#27842)) --- ## 🌐 Web & JavaScript ### WebGPU EP * **Qwen3.5 model support** on WebGPU execution provider ([#27996](microsoft/onnxruntime#27996)) * **QMoE 1-token decode path optimization** — fused operations to reduce GPU dispatches for improved performance ([#27998](microsoft/onnxruntime#27998)) --- ## 🐛 Bug Fixes ### Core Runtime Fixes * **Improved filesystem error messages** during Linux device discovery for better debugging experience ([#27289](microsoft/onnxruntime#27289)) * Fixed missing include for `SetRawDataInTensorProto` in NVIDIA TensorRT RTX tests ([#28065](microsoft/onnxruntime#28065)) --- ## 🙏 Contributors Thanks to our **7 contributors** for this release: [@guschmue](https://github.com/guschmue), [@sanaa-hamel-microsoft](https://github.com/sanaa-hamel-microsoft), [@apsonawane](https://github.com/apsonawane), [@eserscor](https://github.com/eserscor), [@ishwar-raut1](https://github.com/ishwar-raut1), [@qjia7](https://github.com/qjia7), [@theHamsta](https://github.com/theHamsta) **Full Changelog**: microsoft/onnxruntime@v1.25.0...v1.25.1 ## 1.25.0 ## 📢 Announcements & Breaking Changes ### Build & Platform * **C++20 is now required** to build ONNX Runtime from source. Minimum toolchains: MSVC 19.29+, GCC 10+, Clang 10+. Users of prebuilt packages are unaffected. ([#27178](microsoft/onnxruntime#27178)) * **CUDA minimum version raised to 12.0** — CUDA 11.x is no longer supported. Users pinned to CUDA 11.x should stay on ORT 1.24.x or upgrade their CUDA toolkit/driver. ([#27570](microsoft/onnxruntime#27570)) * **ONNX upgraded to 1.21.0** ([#27601](microsoft/onnxruntime#27601)) * **sympy is now an optional dependency** for Python builds. ([#27200](microsoft/onnxruntime#27200)) ### Execution Provider Changes * **ArmNN EP has been removed.** Users should remove any `--use_armnn` build flags and migrate to the MLAS/KleidiAI-backed CPU EP or QNN EP for Qualcomm hardware. ([#27447](microsoft/onnxruntime#27447)) ### API Version * **ORT_API_VERSION** updated to **25**. ([#27280](microsoft/onnxruntime#27280)) --- ## 🔒 Security Fixes * Fixed **potential integer truncation leading to heap out-of-bounds read/write** ([#27544](microsoft/onnxruntime#27544)) * Addressed **Pad Reflect vulnerability** ([#27652](microsoft/onnxruntime#27652)) * **Security fix for transpose optimizer** ([#27555](microsoft/onnxruntime#27555)) * Upgraded minimatch 3.1.2 → 3.1.4 for **CVE-2026-27904** ([#27667](microsoft/onnxruntime#27667)) * Hardened shell command handling for constant strings ([#27840](microsoft/onnxruntime#27840)) * Added validation of `onnx::TensorProto` data size before allocation ([#27547](microsoft/onnxruntime#27547)) * Cleaned up external data path validation ([#27539](microsoft/onnxruntime#27539)) * Fixed misaligned address reads for tensor attributes from raw data buffers ([#27312](microsoft/onnxruntime#27312)) * Fixed **CPU Attention overflow** issue ([#27822](microsoft/onnxruntime#27822)) * Fixed **CPU LRN integer overflow** issues ([#27886](microsoft/onnxruntime#27886)) * Additional input validation hardening: * Tile kernel dim overflow ([#27566](microsoft/onnxruntime#27566)) * Out-of-bounds read in cross entropy ([#27568](microsoft/onnxruntime#27568)) * TreeEnsembleClassifier attributes ([#27571](microsoft/onnxruntime#27571)) * AffineGrid ([#27572](microsoft/onnxruntime#27572)) * EmbedLayerNorm position_ids ([#27573](microsoft/onnxruntime#27573)) * RotaryEmbedding position_ids ([#27597](microsoft/onnxruntime#27597)) * RoiAlign batch_indices ([#27603](microsoft/onnxruntime#27603)) * MaxUnpool indices ([#27432](microsoft/onnxruntime#27432)) * QMoECPU swiglu OOB ([#27748](microsoft/onnxruntime#27748)) * SVMClassifier initializer ([#27699](microsoft/onnxruntime#27699)) * Col2Im SafeInt ([#27625](microsoft/onnxruntime#27625)) --- ## ✨ New Features ### 🔌 Execution Provider Plugin API & CUDA Plugin EP ... (truncated) Commits viewable in [compare view](microsoft/onnxruntime@v1.24.4...v1.30.0). </details> Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore <dependency name> major version` will close this group update PR and stop Dependabot creating any more for the specific dependency's major version (unless you unignore this specific dependency's major version or upgrade to it yourself) - `@dependabot ignore <dependency name> minor version` will close this group update PR and stop Dependabot creating any more for the specific dependency's minor version (unless you unignore this specific dependency's minor version or upgrade to it yourself) - `@dependabot ignore <dependency name>` will close this group update PR and stop Dependabot creating any more for the specific dependency (unless you unignore this specific dependency or upgrade to it yourself) - `@dependabot unignore <dependency name>` will remove all of the ignore conditions of the specified dependency - `@dependabot unignore <dependency name> <ignore condition>` will remove the ignore condition of the specified dependency and ignore conditions </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
This was referenced Sep 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This pull request improves the safety and robustness of memory allocation in the
BFCArenaallocator by adding overflow protection and corresponding tests. The most important changes are:Overflow Protection and Safety:
BFCArena::RoundedBytesfunction inbfc_arena.ccto useSafeInt<size_t>for arithmetic operations, preventing integer overflows during memory rounding calculations.safeint.hheader to enable safe integer operations inbfc_arena.cc.Testing:
RoundedBytesOverflowThrows, inbfc_arena_test.ccto verify that an overflow during allocation throws anOnnxRuntimeException.<limits>header inbfc_arena_test.ccto support boundary value tests.